Papers with French language
French GossipPrompts: Dataset For Prevention of Generating French Gossip Stories By LLMs (2024.eacl-short)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are undergoing a dynamic transformation . however, there is a potential risk of LLMs creating gossips when prompted with contexts . |
| Approach: | a dataset is used to identify prompts that lead to the creation of gossipy content in the french language. |
| Outcome: | a new dataset identifies prompts that lead to the creation of gossipy content in the french language . the model achieves an accuracy of 89.95% . |
Multi-lingual neural title generation for e-Commerce browse pages (N18-3)
Copied to clipboard
| Challenge: | e-Commerce websites are automatically generating millions of browse pages . manual creation of titles is infeasible due to the huge number of browse page types . |
| Approach: | They propose to use sequence-to-sequence models to generate titles for languages . they train the models on multi-lingual data, thereby creating one joint model . |
| Outcome: | The proposed model can generate titles in three different languages, with a focus on low-resource French. |
Providing Semantic Knowledge to a Set of Pictograms for People with Disabilities: a Set of Links between WordNet and Arasaac: Arasaac-WN (2020.lrec-1)
Copied to clipboard
Didier Schwab, Pauline Trial, Céline Vaschalde, Loïc Vial, Emmanuelle Esperanca-Rodier, Benjamin Lecouteux
| Challenge: | Pictograms are a tool that is increasingly used by people with cognitive or communication disabilities. |
| Approach: | They propose a database that links WordNet and Arasaac to link pictograms to semantic knowledge. |
| Outcome: | The proposed database links pictograms with WordNet and Arasaac to create language-independent prototypes. |
A Computational Analysis and Exploration of Linguistic Borrowings in French Rap Lyrics (2024.acl-srw)
Copied to clipboard
| Challenge: | rap is a popular genre in the u.s. and has been used in countries far beyond the uk . linguistic borrowings are especially intriguing in countries such as the eu and europe . |
| Approach: | They manually annotate a lexicon of over 700 borrowings in the French language . they find that there are increases in the proportion of linguistic borrowings, interjections, and Niger-Congo borrowings . |
| Outcome: | The proposed method analyzes a corpus of over 8000 french rap song lyrics and shows that rap borrowings are increasing in prevalence and interjections are decreasing. |
CoFiF Plus: A French Financial Narrative Summarisation Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing corpora for financial narrative summarisation exists in English but there is a significant lack of financial text resources in the French language. |
| Approach: | They propose to use natural language processing to analyse financial documents to find the best summarisation methods. |
| Outcome: | The proposed dataset is the first to provide a comprehensive set of financial text written in French. |
Speech Resources in the Tamasheq Language (2022.lrec-1)
Copied to clipboard
Marcely Zanon Boito, Fethi Bougares, Florentin Barbier, Souhir Gahbiche, Loïc Barrault, Mickael Rouvier, Yannick Estève
| Challenge: | In this paper, we present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger . we share unlabeled audio data in five languages: french, Fulfulde, Hausa, Tamaheq and Zarma . |
| Approach: | They present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger. |
| Outcome: | The proposed datasets are used in the IWSLT 2022 low-resource speech translation track . they consist of radio recordings from daily broadcast news in Niger and Mali . |
FQuAD2.0: French Question Answering and Learning When You Don’t Know (2022.lrec-1)
Copied to clipboard
| Challenge: | Question Answering, including Reading Comprehension, has seen significant scientific breakthroughs over the past few years . but most of these breakthroughs are centered on the English language . |
| Approach: | They propose a dataset to train Question Answering models in the French language . they extend the dataset to 17,000+ unanswerable questions annotated adversarially . |
| Outcome: | The proposed dataset makes it possible to train French Question Answering models with the ability to distinguish unanswerable questions from answerable ones. |
NLP Analytics in Finance with DoRe: A French 250M Tokens Corpus of Corporate Annual Reports (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent advances in neural computing and word embeddings for semantic processing open many new applications areas which had been left unaddressed due to inadequate language understanding capacity. |
| Approach: | They propose a French and dialectal French corpus for NLP analytics in finance, regulation and investment. |
| Outcome: | The proposed corpus is designed to be as modular as possible to allow for maximum reuse in different tasks pertaining to Economics, Finance and Investment. |
Know When to Fuse: Investigating Non-English Hybrid Retrieval in the Legal Domain (2025.coling-main)
Copied to clipboard
| Challenge: | Existing research focuses on a limited set of retrieval methods, evaluated in pairs on domain-general datasets exclusively in English. |
| Approach: | They evaluate the efficacy of hybrid search across a variety of retrieval models in the french language . they find that fusion of different domain-general models consistently enhances performance . |
| Outcome: | The proposed model improves in-domain performance compared to a single model in a zero-shot context . the proposed model also improves when the models are trained in- domain . |
A Benchmark of French ASR Systems Based on Error Severity (2025.coling-main)
Copied to clipboard
| Challenge: | Automatic Speech Recognition (ASR) transcription errors are often assessed using metrics that compare them with a reference transcription. |
| Approach: | They propose to categorize transcription errors into four levels of severity based on objective linguistic criteria, contextual patterns, and the use of content words as the unit of analysis. |
| Outcome: | The proposed evaluation categorizes errors into four levels of severity based on objective linguistic criteria, contextual patterns, and the use of content words as the unit of analysis. |
GeNRe: A French Gender-Neutral Rewriting System Using Collective Nouns (2025.findings-acl)
Copied to clipboard
| Challenge: | Gender rewriting is an NLP task that uses gendered forms to mitigate gender biases. |
| Approach: | They propose a French gender-neutral rewriting system using collective nouns, which are gender-fixed in French. |
| Outcome: | The proposed system detects gendered forms and replaces them with neutral or opposite forms. |
PAGnol: An Extra-Large French Generative Model (2022.lrec-1)
Copied to clipboard
Julien Launay, E.l. Tommasone, Baptiste Pannier, François Boniface, Amélie Chatelain, Alessandro Cappelli, Iacopo Poli, Djamé Seddah
| Challenge: | a growing number of pre-trained language models are available in many different languages. |
| Approach: | They propose a French-language GPT model with scaling laws to train it efficiently . they evaluate the models on discriminative and generative tasks in French . |
| Outcome: | The proposed model trains with the same computational budget as CamemBERT, a model 13 times smaller. |
Introducing RezoJDM16k: a French KnowledgeGraph DataSet for Link Prediction (2022.lrec-1)
Copied to clipboard
Mehdi Mirzapour, Waleed Ragheb, Mohammad Javad Saeedizade, Kevin Cousot, Helene Jacquenet, Lawrence Carbon, Mathieu Lafourcade
| Challenge: | Knowledge graphs are used for information extraction, search engines, question answering, and recommendation systems. |
| Approach: | They propose a French knowledge graph dataset based on RezoJDM. |
| Outcome: | The proposed dataset can be used in many downstream tasks for the French language . it shows that it embeds knowledge graph baselines for link prediction tasks . |
FRACAS: a FRench Annotated Corpus of Attribution relations in newS (2024.lrec-main)
Copied to clipboard
| Challenge: | Quotation extraction is a useful task, but it is not widely studied in other languages. |
| Approach: | They propose to annotate a manually annotated corpus of 1,676 newswire texts in French for quotation extraction and source attribution. |
| Outcome: | The proposed system is compared to the most recent system for quotation extraction in the French language. |
FReND: A French Resource of Negation Data (2024.lrec-main)
Copied to clipboard
| Challenge: | Negation data are limited by the language models of the BERT-generation, which are still underperforming on tasks and benchmarks featuring negation. |
| Approach: | FReND is a freely available corpus of French language in which negations are hand-annotated by their cues and scopes. |
| Outcome: | FReND is the largest dataset available for french negation analysis . it is a valuable resource for linguistic research and as training data for AI tasks such as negation detection. |
French Tweet Corpus for Automatic Stance Detection (2020.lrec-1)
Copied to clipboard
| Challenge: | a new corpus of tweets is being developed for automatic stance detection of fake news . the task involves determining the attitude expressed in a text toward a target . this is a difficult task to overcome as discussions about fake news are controversial . |
| Approach: | They propose to build a human-annotated corpus for automatic stance detection of tweets in french . they propose to use four classes broadly adopted by the community for annotation . |
| Outcome: | The proposed corpus is the first freely available stance annotated tweet corpus in the french language. |
DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains (2023.acl-long)
Copied to clipboard
Yanis Labrak, Adrien Bazoge, Richard Dufour, Mickael Rouvier, Emmanuel Morin, Béatrice Daille, Pierre-Antoine Gourraud
| Challenge: | Recent studies have shown that pre-trained language models improve performance on a wide range of NLP tasks. |
| Approach: | They propose to use pre-trained language models to train medical domains on French language to compare performance with specialized ones. |
| Outcome: | The proposed models can take advantage of existing biomedical models in a foreign language by further pre-training them on our targeted data. |